fix(rewrite): prevent walrus operator double evaluation in assertions - #14447
fix(rewrite): prevent walrus operator double evaluation in assertions#14447RonnyPfannschmidt wants to merge 9 commits into
Conversation
There was a problem hiding this comment.
Pull request overview
Fixes assertion rewriting so walrus (:=) expressions are not evaluated multiple times, preventing side effects from running twice and producing incorrect rewritten-assert behavior (per #14445).
Changes:
- Removes the prior
variables_overwrite/scope-tracking mechanism and adjustsNamedExprhandling to avoid re-evaluation in explanations. - Updates BoolOp/Compare rewriting to stabilize conditions/operands for explanation formatting.
- Adds new regression tests for walrus side-effect/double-evaluation cases and updates existing expected assertion output.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
src/_pytest/assertion/rewrite.py |
Refactors assertion-rewrite AST generation around NamedExpr, BoolOp, and Compare to avoid walrus re-evaluation. |
testing/test_assertrewrite.py |
Updates expected assertion output and adds regression tests for #14445 scenarios. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Pierre-Sassoulas
left a comment
There was a problem hiding this comment.
LGTM, but we probably want another person to look at it
There was a problem hiding this comment.
I haven't reviewed the code yet, before that I dumped the rewritten AST before/after this PR, on the following example:
def side_effect():
return True
def test_walrus_boolop():
assert (x := side_effect())Before
Module(
body=[
Import(
names=[
alias(name='builtins', asname='@py_builtins')]),
Import(
names=[
alias(name='_pytest.assertion.rewrite', asname='@pytest_ar')]),
FunctionDef(
name='side_effect',
args=arguments(),
body=[
Return(
value=Constant(value=True))]),
FunctionDef(
name='test_walrus_boolop',
args=arguments(),
body=[
If(
test=UnaryOp(
op=Not(),
operand=NamedExpr(
target=Name(id='x', ctx=Store()),
value=Call(
func=Name(id='side_effect', ctx=Load())))),
body=[
Assign(
targets=[
Name(id='@py_format1', ctx=Store())],
value=BinOp(
left=BinOp(
left=Constant(value=''),
op=Add(),
right=Constant(value='assert %(py0)s')),
op=Mod(),
right=Dict(
keys=[
Constant(value='py0')],
values=[
IfExp(
test=BoolOp(
op=Or(),
values=[
Compare(
left=Constant(value='x'),
ops=[
In()],
comparators=[
Call(
func=Attribute(
value=Name(id='@py_builtins', ctx=Load()),
attr='locals',
ctx=Load()))]),
Call(
func=Attribute(
value=Name(id='@pytest_ar', ctx=Load()),
attr='_should_repr_global_name',
ctx=Load()),
args=[
NamedExpr(
target=Name(id='x', ctx=Store()),
value=Call(
func=Name(id='side_effect', ctx=Load())))])]),
body=Call(
func=Attribute(
value=Name(id='@pytest_ar', ctx=Load()),
attr='_saferepr',
ctx=Load()),
args=[
NamedExpr(
target=Name(id='x', ctx=Store()),
value=Call(
func=Name(id='side_effect', ctx=Load())))]),
orelse=Constant(value='x'))]))),
Raise(
exc=Call(
func=Name(id='AssertionError', ctx=Load()),
args=[
Call(
func=Attribute(
value=Name(id='@pytest_ar', ctx=Load()),
attr='_format_explanation',
ctx=Load()),
args=[
Name(id='@py_format1', ctx=Load())])]))])])])After
Module(
body=[
Import(
names=[
alias(name='builtins', asname='@py_builtins')]),
Import(
names=[
alias(name='_pytest.assertion.rewrite', asname='@pytest_ar')]),
FunctionDef(
name='side_effect',
args=arguments(),
body=[
Return(
value=Constant(value=True))]),
FunctionDef(
name='test_walrus_boolop',
args=arguments(),
body=[
If(
test=UnaryOp(
op=Not(),
operand=NamedExpr(
target=Name(id='x', ctx=Store()),
value=Call(
func=Name(id='side_effect', ctx=Load())))),
body=[
Assign(
targets=[
Name(id='@py_format1', ctx=Store())],
value=BinOp(
left=BinOp(
left=Constant(value=''),
op=Add(),
right=Constant(value='assert %(py0)s')),
op=Mod(),
right=Dict(
keys=[
Constant(value='py0')],
values=[
IfExp(
test=BoolOp(
op=Or(),
values=[
Compare(
left=Constant(value='x'),
ops=[
In()],
comparators=[
Call(
func=Attribute(
value=Name(id='@py_builtins', ctx=Load()),
attr='locals',
ctx=Load()))]),
Call(
func=Attribute(
value=Name(id='@pytest_ar', ctx=Load()),
attr='_should_repr_global_name',
ctx=Load()),
args=[
Name(id='x', ctx=Load())])]),
body=Call(
func=Attribute(
value=Name(id='@pytest_ar', ctx=Load()),
attr='_saferepr',
ctx=Load()),
args=[
Name(id='x', ctx=Load())]),
orelse=Constant(value='x'))]))),
Raise(
exc=Call(
func=Name(id='AssertionError', ctx=Load()),
args=[
Call(
func=Attribute(
value=Name(id='@pytest_ar', ctx=Load()),
attr='_format_explanation',
ctx=Load()),
args=[
Name(id='@py_format1', ctx=Load())])]))])])])The diff is:
@@ -57,20 +57,14 @@
attr='_should_repr_global_name',
ctx=Load()),
args=[
- NamedExpr(
- target=Name(id='x', ctx=Store()),
- value=Call(
- func=Name(id='side_effect', ctx=Load())))])]),
+ Name(id='x', ctx=Load())])]),
body=Call(
func=Attribute(
value=Name(id='@pytest_ar', ctx=Load()),
attr='_saferepr',
ctx=Load()),
args=[
- NamedExpr(
- target=Name(id='x', ctx=Store()),
- value=Call(
- func=Name(id='side_effect', ctx=Load())))]),
+ Name(id='x', ctx=Load())]),
orelse=Constant(value='x'))]))),
Raise(
exc=Call(This looks good for the issue, since now we no longer run side_effect twice three times.
However if I tweak in this way:
def side_effect():
return True
def test_walrus_boolop():
assert (x := side_effect()) and (x := False)the assertion is
x.py:5: in test_walrus_boolop
assert (x := side_effect()) and (x := False)
E assert (False and False)
which is incorrect (should be assert (True and False)). That said, this also happens in main.
Let me know if you want to tackle this problem in this PR as well, in which I'll wait before reviewing, or if I should open a separate issue for that and review this PR as is.
|
good find, i'll address it in here |
|
i found a interesting issue about very duplicate tracking, investigating now |
|
now the change is a litte bigger than intended |
acd9277 to
c4369d0
Compare
|
Thanks for working on this. A differential check found three remaining def identity(value):
return value
def test_compare_preserves_pre_walrus_left_value():
value = "Hello"
assert value != identity(value := value.lower())
assert value == "hello"Python evaluates the left operand before the call, so this compares def collect(*values):
return values
def test_call_preserves_earlier_positional_argument():
value = "Hello"
assert collect(value, identity(value := value.lower())) == (
"Hello",
"hello",
)
assert value == "hello"The first positional argument should already be A related false-comparison case is: def test_failed_compare_uses_pre_walrus_left_value():
value = 2
try:
assert value == identity(value := 3)
except AssertionError:
pass
else:
raise AssertionError("assertion was rewritten as 3 == 3")
assert value == 3This must compare Values evaluated before a later walrus need to be captured before that
I can provide tests adapted to These cases were identified with automated assistance, then reduced and |
c4369d0 to
9fe133f
Compare
|
Rebased and extended. Two changes worth calling out. The reported evaluation-order cases are fixed here. @scapalive — thank you, all three reproduce and all three are now covered. They are not regressions from this PR (they fail on The mechanism was already in the PR, just too shallow, and inconsistently so between two visitors this PR touches: This now sits on #14813, a test-only PR that adds a coverage matrix for the rewriter and records every known gap as a strict xfail. This PR closes four of those groups — One |
9fe133f to
6caf74a
Compare
6caf74a to
6c9f83e
Compare
6c9f83e to
fa8ff7a
Compare
Pierre-Sassoulas
left a comment
There was a problem hiding this comment.
I didn't see anything shocking by skimming. i'll review in details later.
Maybe the helper script to compare two pytest versions' output could be in their own PR ? In pylint and mypy there is a primer that permits to see the change in output in a feature branch compared to main by running pylint/mypy on a choice selection of open source repos. Maybe it could be adapted for pytest based on those two scripts.
Also the big file with separator as comment could be burst into a directory of multiple files?
|
It would be nice to see original/rewritten AST for some simple case, to get a quick sense of the new method, before diving into the code. If the AST can also be written in source code form it would make it even easier. Maybe for the example given above: def side_effect():
return True
def test_walrus_boolop():
assert (x := side_effect()) and (x := False)Sorry I'm too lazy to do it myself... |
| assert result.ret == 0 | ||
|
|
||
|
|
||
| class TestIssue14445: |
There was a problem hiding this comment.
Are these tests redundant with the coverage ones, or are they testing something separate?
Also in #14813 (comment) you said you'll delete TestAssertionRewriteWalrusOperator here. Is it not redundant now?
dead644 to
5fe5987
Compare
|
To bluetech comment, I made a pytest plugin to do golden master / caracterisation tests (pytest-remaster) which is what we want to do here for easy review and update of ast changes imo. Let me know what you think. |
b3f3947 to
0253180
Compare
Fixes pytest-dev#14445 - assertion rewriting evaluated NamedExpr (:=) expressions multiple times, causing side effects to fire repeatedly. The root cause was the `variables_overwrite` mechanism which stored and re-evaluated NamedExpr AST nodes in subsequent assertions, in `_call_reprcompare`'s results tuple, and in explanation formatting. The fix: - visit_NamedExpr: reference the target variable in explanations instead of re-evaluating the full expression - visit_Compare: assign left-side NamedExpr to a temp before right-side hoisting; freeze left_res when a comparator walrus targets the same name; replace NamedExpr entries in `results` with target variables - visit_BoolOp: capture short-circuit condition in a stable temp for the explanation path; remove walrus target rename logic - visit_Call: remove variables_overwrite substitution (walrus now properly assigns to user variables in its natural evaluation position) - Remove variables_overwrite, scope tracking, Sentinel class Co-authored-by: Cursor AI <ai@cursor.sh> Co-authored-by: Anthropic Claude Sonnet 4 <claude@anthropic.com>
Co-authored-by: Cursor AI <ai@cursor.sh> Co-authored-by: Anthropic Claude Sonnet 4 <claude@anthropic.com>
Add tests for two remaining walrus double-evaluation scenarios: - Bare NamedExpr as BoolOp operand evaluated twice via condition check - Same walrus target in chained comparison evaluated multiple times Co-authored-by: Cursor AI <ai@cursor.sh> Co-authored-by: Anthropic Claude Sonnet 4 <claude@anthropic.com>
Use the already-assigned res_var to build the short-circuit condition instead of the raw visitor result, preventing bare NamedExpr operands from being evaluated a second time when checking truthiness. Co-authored-by: Cursor AI <ai@cursor.sh> Co-authored-by: Anthropic Claude Sonnet 4 <claude@anthropic.com>
In a chained comparison like `(x := f()) < (x := g()) < (x := h())`, each NamedExpr comparator is now assigned to a temp variable so it evaluates exactly once. Previously the raw NamedExpr node would be reused as left_res in the next iteration, causing double evaluation. Co-authored-by: Cursor AI <ai@cursor.sh> Co-authored-by: Anthropic Claude Sonnet 4 <claude@anthropic.com>
When multiple walrus operators target the same variable in a BoolOp (e.g., `assert (x := side_effect()) and (x := False)`), the assertion explanation previously showed the final value of `x` for all operands because the format context evaluated lazily after all operands ran. Fix by tracking Name/NamedExpr operand values in stable @py_assert variables (via self.assign) immediately after evaluation, then pointing the explanation format context at the tracked copy. This uses the same value-tracking mechanism already used by visit_Call, visit_Attribute, etc. Fixes the case reported by @bluetech in PR review. Co-authored-by: Cursor AI <ai@cursor.sh> Co-authored-by: Anthropic Claude Sonnet 4 <claude@anthropic.com>
Replace the blanket snapshot-all-operands approach with a targeted one: pre-scan the BoolOp to find walrus targets, then only snapshot operands whose value a later walrus would corrupt. Snapshot rules: - NamedExpr (non-last): always, to avoid re-evaluating side effects - Name with later walrus conflict: to freeze the pre-overwrite value - Everything else: use res directly (stable @py_assert or plain name) Non-walrus BoolOps now generate identical code to 8.3.5 (no snapshots). Co-authored-by: Cursor AI <ai@cursor.sh> Co-authored-by: Anthropic Claude Opus 4 <claude@anthropic.com>
The rewriter hoists each operand into its own statement, but a plain
name is left as a bare load evaluated when the enclosing expression is
assembled -- after the statements of the operands that follow it. A
walrus operator in a later operand rebinds the name in between, so both
the value used and the value reported were the post-walrus one, while
Python evaluates the earlier operand first:
assert value != identity(value := value.lower())
visit_BoolOp already guarded against this; extract its pre-scan as
_walrus_targets() and add visit_operand() to apply the same freeze in
visit_Compare, visit_Call and visit_BinOp. visit_Compare previously
matched only a comparator that *was* a NamedExpr, missing walrus
operators nested inside it; visit_Call did not guard at all, so an
earlier argument saw a later argument's assignment.
These cases predate the walrus rework -- they fail on main too.
Closes the single-eval-walrus, order-compare-left, order-call-argument
and order-binop-left groups in the coverage matrix. order-call-argument
keeps one entry: a bare walrus argument is still substituted into a
later one, which visit_operand does not yet see because the operand is a
NamedExpr rather than a Name.
Reported-by: Denis Scapin
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TestAssertionRewriteWalrusOperator predates the coverage matrix and asks its questions through runpytest(): twelve tests that mostly check ret == 0. Nine of them the matrix already answers in-process and more precisely -- what the failure message says, what the operands saw, how often each ran. Two are worth more than a deletion, so they move rather than vanish: assert not (a and ((a := False) is False)) reads back as an introspection case, and the composite chain assert a and True and ((a := False) is False) and (a is False) and ... as an evaluation-order one. That second is why the deletion is not just tidying: it passes on main, and it passes there because of the bug. Main rewrites the walrus target to an internal temp and substitutes the stored NamedExpr into every later read of the name -- including the next statement, so the `assert a is None` that is supposed to verify the outcome re-runs the walrus and creates it. Asked through returned values instead of ret == 0, main fails the case. TestIssue14445 loses the four tests that restate matrix single-evaluation entries this PR already adds, and keeps the two reproducers from the issue. One test survives in place: nothing the rewriter stores may outlive a statement, which needs two tests in one module to observe and so cannot be said in-process. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
0253180 to
0de28a6
Compare
| @@ -0,0 +1,3 @@ | |||
| Fixed assertion rewriting evaluating walrus operator (``:=``) expressions multiple times, causing incorrect test results when the expression had side effects (e.g., incrementing a counter or calling a function). | |||
Summary
Fixes #14445 — assertion rewriting evaluated walrus operator (
:=) expressionsmultiple times, so an assertion over an expression with side effects reported a
result the test never actually computed.
Root cause: the
variables_overwritemechanism storedNamedExprAST nodesand substituted them back into every later read of the target name — in
subsequent assertions, in
_call_reprcompare's results tuple, and inexplanation formatting. Each substitution is another evaluation.
Fix: remove
variables_overwriteentirely, and insteadorder is the interpreter's, not the rewriter's;
assign()when a later walrus targets thesame name, so the explanation reports what that operand actually saw;
The
self.scope/_SCOPE_END_MARKERbookkeeping goes with it: it existed onlyto key
variables_overwriteper scope.What is in the diff
Four files, ~160 lines of rewriter change:
src/_pytest/assertion/rewrite.pytesting/test_assertrewrite_coverage.pytesting/test_assertrewrite.pychangelog/14445.bugfix.rstCoverage
The new cases in the matrix cover single evaluation (
compare,boolop,chained compare) and evaluation order — that an operand evaluated before a
later walrus is reported with its pre-assignment value, for compare left
operands, call arguments (positional, keyword and
**), BinOp left operands andchained-compare operands.
The three evaluation-order cases @scapalive reported are among them. They are
not regressions from this PR — they fail on
maintoo — but they are the samedefect, so they are fixed and covered here rather than deferred.
Test cull
TestAssertionRewriteWalrusOperatoris gone. Nine of its twelverunpytest()tests are answered more precisely by the matrix, two moved there (one as
introspection, one as evaluation order), and one stayed in
test_assertrewrite.pybecause it needs two tests in one module to sayanything.
TestIssue14445keeps the two reproducers from the issue itself.The composite-chain case is the reason this is a cull and not just tidying: it
passes on
main, and it passes because of the bug.mainrewrites the walrustarget to an internal temp and substitutes the stored
NamedExprinto everylater read of the name, including the next statement — so the
assert a is Nonemeant to verify the outcome re-runs the walrus and produces the value it is
checking for. Asked through returned values instead of
ret == 0,mainfailsthat case.
Review notes
@bluetech — the BoolOp explanation you flagged is fixed on this branch. On
mainreportsassert (False and False); this branch reportswhich is the case now pinned by
test_walrus_in_boolop_reports_each_operand.The AST before/after you asked for is what
scripts/assert_rewrite_diff.pyprints — that script landed separately in #14921, and #14921 shows its output
for exactly this snippet on
main.Status
Rebased onto
main. #14813 (the coverage matrix) and #14921 (the rewrite-diffscript) have both merged, so this branch no longer carries either; earlier
revisions of this description referred to them as pending.
Full suite green on the rebased branch: 4511 passed, 46 skipped, 17 xfailed.
AI-assisted, reviewed and driven by me; see the
Co-authored-bytrailers on thecommits.